Back

Genetic Epidemiology

Wiley

Preprints posted in the last 7 days, ranked by how well they match Genetic Epidemiology's content profile, based on 55 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.1%
9.5%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

2
ICONIC: An R Package for Integrating Instrumental Variable- and Negative-Control-Informed Causal Discovery and Diagnostics in Multiomic Studies

Bresnahan, S. T.; Xiong, C.; Head, T.; Chang, Y.-H.; Bhattacharya, A.; Huang, J. Y.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.26.26361466 medRxiv
Top 0.1%
7.9%
Show abstract

Unmeasured confounding threatens causal inference and replicability in observational multi-omic studies across variable environments. Genetic instrumental variables (Mendelian randomization) and negative-control calibration each address complementary sources of unmeasured confounding, yet no existing framework unifies them for omics-scale mediation analysis. We introduce ICONIC, an R package that embeds genetic instruments and negative controls within a proximal causal inference framework for total-effect and mediation analysis. ICONIC implements eight estimators spanning five confounding-control strategies, supports continuous, binary, and time-to-event outcomes, and provides extensive diagnostics including sensitivity analyses that map estimator performance across plausible assumptions. Ground-truth benchmarks are calibrated to real-omics covariance structures via a hybrid generative model (GAN + feature-level Gaussian copula) rather than parametric simulation, and a companion planning tool predicts performance gains from collecting additional omic data. We demonstrate ICONIC in two case studies: identifying placental transcriptomic mediators of gestational diabetes on birth weight (n = 164), and tumor-expression mediators of smoking intensity on lung cancer survival (n = 494). Notably, ICONIC's diagnostics recommended different estimation strategies across the two scenarios, reflecting differences in the likely influence of unmeasured confounding. ICONIC is freely available at https://github.com/sbresnahan/iconic/.

3
MOSurvivor-Guided Joint CpG Selection and XGBoost Hyperparameter Optimization for Compact Epigenetic Age Prediction

Yelgi, A.; Tavangari, S.; Shakarami, Z.; Janfaza, S.

2026-08-29 genomics 10.64898/2026.08.26.747213 medRxiv
Top 0.2%
2.5%
Show abstract

Accurate epigenetic age prediction from DNA methylation profiles is intrinsically high-dimensional, creating a need for parsimonious models that preserve predictive performance while reducing the number of assayed cytosine-phosphate-guanine (CpG) loci. This study introduces MOSurvivor, a population-based multi-objective search framework that jointly optimizes a weight-threshold CpG selector and eight XGBoost hyperparameters. Experiments used the GSE40279 whole-blood cohort (656 individuals profiled on the Illumina HumanMethylation450 platform). After retaining 1,000 age-correlated CpGs, five strategies were evaluated on the same 30 seeded 80:20 train/test splits: fixed-parameter XGBoost using all 1,000 CpGs, random search, a genetic algorithm, particle swarm optimization, and MOSurvivor. Internal fitness was estimated using three-fold cross-validation on each training set. Across the 30 held-out test sets, MOSurvivor achieved a mean absolute error (MAE) of 4.149 {+/-} 0.300 years, root mean squared error of 5.545 {+/-} 0.392 years, and R2 of 0.855{+/-} 0.027 while retaining 211.6 {+/-} 54.8 CpGs. Relative to full-feature XGBoost (MAE 4.095 {+/-} 0.285 years), MOSurvivor reduced the feature set by 78.8% at an MAE increase of only 0.054 years (1.3%). Paired Wilcoxon tests found no significant accuracy difference between MOSurvivor and any comparator (all unadjusted p > 0.05; all Holm-adjusted p [≥] 0.476). The most recurrent locus, cg16867657, appeared in 29 runs, whereas mean pairwise Jaccard similarity was 0.124, indicating a small stable core embedded in multiple near-equivalent feature subsets. MOSurvivor thus offers a competitive accuracy-parsimony trade-off rather than superior absolute accuracy. External validation and leakage-free nested feature preselection remain necessary before biological or clinical translation. Keywords: epigenetic clock, DNA methylation, feature selection, multi-objective optimization, XGBoost, metaheuristics, biological aging.

4
Addressing Measurement Error of Machine-Learned Physical Activity in Nonlinear Dose-Response Survival Analysis: Development and Evaluation of Accelerated Failure Time, Spline, and Simulation-Extrapolation Method

Mamiya, H.; Zhang, Q.; Zhang, X.; Yan, Y.; Sharma, A.

2026-08-31 epidemiology 10.64898/2026.08.25.26361155 medRxiv
Top 0.4%
1.5%
Show abstract

Wearable (accelerometer) data and machine-learning allow objective assessment of the amount of daily physical activity. However, wearable-derived human activity is subject to measurement error. No studies have corrected the dose-response association between physical activity and survival time to chronic diseases, including cardiovascular disease (CVD). The objective is to estimate the measurement error-corrected association between CVD events and multiple measures of daily duration of light and total physical activity, derived from machine-learning and conventional accelerometer-processing methods. Our method combined an accelerated failure time model, spline, and simulation-extrapolation (SIMEX). The method recovered the true dose-response non-linear association in simulated data, while the naive model failed to capture it due to substantial attenuation. Application to the UK Biobank accelerometer cohort also showed an increased protective association of total physical activity after SIMEX correction (Time Ratio [TR] = 1.56, 95% CI: 1.28-1.82 vs. TR = 1.38, 95% CI: 1.24-1.54 for SIMEX-corrected vs. uncorrected dose-response association between the 95th and 5th percentiles of total activity), with a similar increase for light physical activity. Sensitivity analysis indicates that the female population experiences a substantially larger protective association after SIMEX correction than males. Dose-response survival analysis is a widely used analytical method in physical activity epidemiology and benefits from measurement error correction.

5
Hybrid risk scores integrating polygenic and clinical variables for endometriosis prediction

Goroshchuk, O.; Koller, D.

2026-09-03 epidemiology 10.64898/2026.08.31.26361798 medRxiv
Top 0.5%
1.1%
Show abstract

Background: Endometriosis affects approximately 10% of reproductive-age women and is associated with substantial diagnostic delay and heterogeneous symptom presentation. Prior machine-learning prediction models have relied on comorbidity data alone or on small candidate-variant genetic scores, with inconsistent or incompletely reported performance. No study has combined a well-powered, multi-ancestry polygenic risk score (PRS) with environmental, reproductive, and symptom data in a single hybrid model. We developed and evaluated hybrid risk-prediction models integrating a genome-wide, multi-ancestry PRS with clinical and symptom data for endometriosis in the US-based All of Us Research Program. Methods: Among 69,376 participants (15,382 endometriosis cases, 53,994 controls) across six genetically inferred ancestry groups, we computed individual-level PRS values using PRS-CS weights derived from an independent, multi-ancestry GWAS. Five nested logistic regression, random forest, and XGBoost models progressively added age, ancestry, and within-ancestry genetic principal components (Model 1), environmental and reproductive factors (Model 2), symptom and comorbidity indicators (Model 3), all covariates combined (Model 4), and PRS x environment interactions (Model 5). Performance was assessed by AUROC in a held-out test set and 5-fold cross-validation, with class-weighted, Youden-optimized thresholds used for sensitivity, specificity, and predictive values; permutation importance identified top contributors. Pairwise AUROC differences were tested with a Holm-corrected DeLong-type test. Results: Discrimination improved from AUROC 0.63 (PRS, age, ancestry, principal components) to 0.72 for the full model, driven mainly by symptom and comorbidity data. XGBoost consistently outperformed logistic regression and random forest. The PRS ranked among the top individual predictors by permutation importance in nearly every model, alongside age, while genetic and demographic information alone gave only modest discrimination, and PRS x environment interactions did not improve on environmental factors alone. Threshold optimization yielded balanced sensitivity and specificity (~0.67/0.65) versus near-zero sensitivity at a default threshold. Conclusions: Combining the PRS with symptom and comorbidity data gave the best discrimination compared to solely a well-powered, multi-ancestry PRS as a predictor of endometriosis. This study clarifies both the promise and current limits of hybrid genetic-clinical prediction for endometriosis and points to symptom-based phenotyping, molecular subtyping, and external validation as priorities.

6
The Role of Distress-related Metabolic Dysfunction in Ovarian Cancer Development: a pooled case-control study

Lin, N.; Balasubramanian, R.; Menichetti, G.; Eliassen, H.; Trabert, B.; Avila-Pacheco, J.; Townsend, M. K.; Terry, K. L.; Clish, C. B.; Tworoger, S. S.; Zeleznik, O. A.

2026-08-31 epidemiology 10.64898/2026.08.27.26361473 medRxiv
Top 0.5%
1.1%
Show abstract

Background: Evidence suggests chronic distress influences ovarian cancer (OC) etiology and metabolomic profiles. Here, we evaluated the association of a metabolite-based distress score (MDS) and OC risk. Methods: We included two matched case-control studies nested within the Nurses' Health Studies (N=584) and the Prostate, Lung, Colorectal, and Ovarian Cancer Screening Trial (N=348). Metabolites were measured 3-27 years before diagnosis using liquid-chromatography tandem mass spectrometry. We examined the association of quintiles of MDS and 19 constituent metabolites with OC risk using unconditional logistic regression and stratified by tumor histotype, menopausal status, and age at diagnosis. Results: We observed women in the highest versus lowest quintile of MDS had an increased OC risk (OR=1.62,95%CI=1.03-2.54,ptrend=0.07), and type 2 tumors (OR=1.71,95%CI=1.03-2.83,ptrend=0.11). Associations were suggestively stronger for premenopausal and <69-year-old women, and driven by pseudouridine, and N2,N2-dimethylguanosine. Conclusion: Our findings suggest chronic distress-associated metabolic dysregulation may represent a novel OC risk factor, especially among younger women.

7
Investigating adiposity in childhood and adulthood on later life sleep health: a lifecourse Mendelian randomization study

pathak, s.; Richardson, T.; Sanderson, E.; Arora, N.; Strand, L.; Asvold, B. O.; Bhatta, L.; Brumpton, B.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.27.26361310 medRxiv
Top 0.7%
0.6%
Show abstract

Background: Higher Body Mass Index (BMI) is an established risk factor of sleep disturbance. It is not known if the effect is homogeneous across the lifecourse or if there is a particular time point in life that might be best to target. Methods: Two-sample Mendelian randomization (MR) was used to investigated the effect of childhood adiposity (adjusting on adulthood adiposity and obstructive sleep apnea (OSA)) on insomnia, morning chronotype, sleep duration, daytime sleepiness and daytime napping. Similarly, total, and direct effect of adulthood adiposity on these outcomes was explored. We used summary statistics from a genome-wide association study (GWAS) of UK Biobank for childhood and adulthood adiposity (n=453,169) and large-scale consortia of OSA (Million Veteran Program) (n=410,268), insomnia, and chronotype (23andMe) (n=1,978,022 and n=248,1000, respectively). Results: Two-sample univariable MR analysis provided no evidence of an effect of genetically predicted childhood adiposity on later life insomnia (Odds ratio (OR)= 0.94, 95% Confidence interval (CI)= 0.87, 1.03). Whereas, multivariable MR (adjusted for adulthood adiposity) analysis provide strong evidence of direct protective effect of genetically predicted childhood adiposity on later life insomnia (OR= 0.70, CI= 0.64, 0.77). Further, both in univariable and multivariable MR, a strong positive effect of increased childhood body size on morning chronotype was observed (OR= 1.16, CI= 1.01, 1.33 and OR= 1.36, CI= 1.15, 1.62, respectively) after accounting for adulthood body size. In both analysis the estimate did not change considerably after aditionally adjusting for OSA. However, childhood and adulthood adiposity found to be associated with OSA and OSA with insomnia. In both univariable and multivariable analysis, increased body size in adulthood increased the risk of having insomnia and a morning chronotype. Conclusions: The findings suggest that higher body size in childhood is not a risk factor for later life insomnia, whereas higher body size in adulthood was. Further, if healthy body size is maintained in adulthood, high childhood adiposity may decrease the risk of insomnia and increase the risk of being a morning person in later life. Keywords: childhood, adulthood, obesity, insomnia, morning chronotype, medelian randomization

8
Artificial Scientific Intelligence for Measurement-burden-aware Modelling and Interpretation of Multi-site Bone Mineral Density

Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.

2026-09-01 health informatics 10.64898/2026.08.30.26361665 medRxiv
Top 0.8%
0.6%
Show abstract

Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.

9
Publication Bias in Abstracts Presented at the American Diabetes Association Scientific Sessions: A Retrospective Cohort Study

Pinedo-Torres, I.; Taype-Rondan, A.; Vera-Luza, A. A.; Zegarra-Lizana, P. A.; Rojas-Vilca, J. L.; Yovera-Aldana, M.

2026-08-31 epidemiology 10.64898/2026.08.26.26361486 medRxiv
Top 1%
0.4%
Show abstract

Objective. To determine the publication rate of abstracts presented at the American Diabetes Association Scientific Sessions and to evaluate the association between statistical significance of study results and subsequent publication. Research Design and Methods. We conducted a retrospective cohort study of abstracts presented at the 2018 American Diabetes Association Scientific Sessions. The primary exposure was study result category (statistically significant vs. non-statistically significant findings), and the primary outcome was publication in an indexed journal within 5 years after conference presentation. Publication status was determined through PubMed/MEDLINE and Scopus searches. Adjusted relative risks (RRs) and 95% CIs were estimated using generalized linear models with Poisson distribution and robust variance. Results. Among 541 included abstracts, 321 (59.3%) were subsequently published in indexed journals. Abstracts reporting statistically significant findings had a higher publication rate than those reporting non-statistically significant findings (61.9% vs. 42.3%; p=0.002). In the adjusted analysis, abstracts with non-statistically significant findings had a lower likelihood of publication compared with those reporting statistically significant findings (adjusted RR 0.71 [95% CI 0.55-0.93]; p=0.013). Conclusions. Approximately four in ten abstracts presented at the ADA Scientific Sessions were not published within 5 years. Abstracts reporting non-statistically significant findings had a lower likelihood of subsequent publication, suggesting persistent publication bias in diabetology research. Future initiatives promoting the interpretation of effect estimates, confidence intervals and clinical relevance, rather than statistical significance alone, may help reduce selective dissemination of evidence

10
A novel framework leveraging non-causal associations reveals shared pathways linking inflammation and cancer risk

Yarmolinsky, J.; Cavallo, F. R.; Koskeridis, F.; Yu, X.; Bouras, E.; Richenberg, G.; Costantini, I.; Ray, D.; Woolf, B.; Karhunen, V.; Ellis, L.; Haycock, P. C.; Hemani, G.; Davey Smith, G.; Tsilidis, K. K.; Zuber, V.; McKay, J. D.; Dehghan, A.; Tzoulaki, I.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.30.26361622 medRxiv
Top 1%
0.3%
Show abstract

Confounding is a central challenge in observational studies. Here, we propose a framework for identifying confounders of two non-causally related traits by employing cross-trait pleiotropy analysis to detect genetic loci that affect both traits and multi-trait colocalisation to identify molecular phenotypes mediating these effects. We apply this approach to the analysis of C-reactive protein (CRP) - a non-specific marker of inflammation - and 10 inflammation-related cancers. In UK Biobank, higher pre-diagnostic CRP levels are associated with increased risk of multiple cancers, but bidirectional Mendelian randomization provides little evidence for a causal relationship. Cross-trait genetic analyses identify 92 loci with shared CRP-cancer effects including those with established roles in cancer and 50 novel loci such as RSPO3 (breast cancer) and GCKR (colorectal cancer). Integration with proteomic and single-cell transcriptomic data identified putative molecular mediators at 24 loci including plasma TLR1 levels in breast cancer and CD4+ T cell IRF5 expression in kidney cancer. Notably, 15 candidate effector genes encode targets of approved or investigational medications, including IL6, PDE4D, and CASP8, indicating potential opportunities for their repurposing for cancer prevention. The proposed approach provides a generalisable framework for leveraging non-causal phenotypic relationships to yield insights into disease mechanisms and therapeutic targets for disease prevention.

11
A Randomized Non-Inferiority Trial of an eHealth Delivery Alternative for Cancer Genetic Testing for Hereditary Cancer (eREACH2)

Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.

2026-09-03 genetic and genomic medicine 10.64898/2026.09.01.26361920 medRxiv
Top 2%
0.2%
Show abstract

Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.

12
Socioeconomic position, adverse childhood experiences, and menstrual symptoms in two generations of a prospective UK cohort.

Sawyer, G.; Farooq, B.; Birnie, K.; Fraser, A.; Lawlor, D. A.; Sharp, G. C.; Howe, L. D.

2026-08-31 epidemiology 10.64898/2026.08.27.26361513 medRxiv
Top 2%
0.1%
Show abstract

Background: Inequalities exist for many health outcomes, but there is limited evidence regarding menstrual symptoms despite their importance for health and wellbeing. We aimed to investigate inequalities in menstrual symptoms according to socioeconomic position and childhood adversity. Methods: In two generations (G0 mothers and G1 offspring) from the Avon Longitudinal Study of Parents and Children (ALSPAC), a UK prospective cohort study, we examined associations of multiple indicators of socioeconomic position (SEP) and adverse childhood experiences (ACEs) with menstrual symptoms (pain, abnormal uterine bleeding, and premenstrual syndrome (PMS) measured 3-8-years post-birth in G0 and 17-21-years-old in G1), using multivariable logistic regression. Samples ranged from 4,828 to 9,335 G0 participants and 1,288 to 2,757 G1 participants depending on the exposure-outcome association. Missing data were addressed using multiple imputation and inverse probability weighting. Results: Financial difficulties were associated with greater odds of menstrual pain (G1 OR 1.41; 95% CI 1.07, 1.86: G0 OR 1.55; 95% CI 1.36, 1.76) and irregular cycles (G1 OR 1.60; 95% CI 1.12, 2.29: G0 OR 1.48; 95% CI 1.27, 1.72) in both generations, as well as with short/long cycle lengths in G0 only. Lower education and manual social class were also associated with these three menstrual symptoms in at least one generation. Conversely, higher SEP was associated with PMS in both generations. Higher cumulative ACEs were consistently associated with menstrual pain (4+ compared to none: G1 OR 2.15; 95% CI 1.48, 3.11: G0 OR 1.52; 95% CI 1.29, 1.80) and irregular cycles (G1 OR 1.92; 95% CI 1.20, 3.09: G0 OR 1.54; 95% CI 1.26, 1.87) but not cycle length. Lower parental education, financial difficulties, and cumulative ACEs were associated with heavy bleeding in G1 offspring only, whereas financial difficulties, own manual social class, and cumulative ACEs were associated with prolonged bleeding in G0 mothers only. Higher cumulative ACEs were also associated with PMS in G1 offspring only. Conclusions: We found evidence of inequalities according to socioeconomic disadvantage and childhood adversity for multiple menstrual symptoms, although some associations were only observed in one generation. Findings suggest that menstrual symptoms are disproportionately experienced by socially and socioeconomically disadvantaged women.

13
SALRR: Scalable Analysis of Long-Read RNA-Seq Enables Comprehensive Transcriptome Profiling in Human Brain

Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.

2026-08-29 genomics 10.64898/2026.08.27.747499 medRxiv
Top 2%
0.1%
Show abstract

Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.

14
Developing the Longitudinal Study of Aging in Guatemala (ELEGUA): Rationale and pilot protocol

Corzantes, K.; Choy, K.; Adar, S.; Castellanos, L. F.; Gross, A. L.; Langa, K. M.; Rohloff, P.; Weerman, B.; Briceno, E.; Ramirez-Zea, M.; Behrman, J.; Flood, D.

2026-08-31 epidemiology 10.64898/2026.08.26.26361136 medRxiv
Top 2%
0.1%
Show abstract

Introduction Guatemala is the most populous country in Central America and a setting with unique opportunities for aging research. Approximately 40% of Guatemala's population is Indigenous Maya, who together speak 22 Mayan languages. Currently, there is no population-based aging study in Guatemala and few aging studies in Latin America among Indigenous populations. The Longitudinal Study of Aging in Guatemala (ELEGUA) aims to address these gaps by developing a nationally representative, population-based, longitudinal aging study modeled on the Health and Retirement Study and the Harmonized Cognitive Assessment Protocol, adapted to the cultural and linguistic context of Guatemala. The objective of this protocol is to describe the rationale and design of the ELEGUA pilot survey. Methods and analysis The ELEGUA pilot was a cross-sectional household survey of adults aged 40 years or older in Tecpan, Guatemala. Tecpan was chosen because its diverse population facilitated testing of study procedures in both Spanish and Kaqchikel, a common Mayan language. The survey included up to 600 households sampled using a multistage stratified cluster design. Within each household, one individual aged 40 years or older was selected, with oversampling of adults aged 55 years or older. This respondent completed a comprehensive questionnaire, including detailed cognitive tests, and provided physical measurements and a venous blood sample. Household respondents provided information on household economics and family structure, and an informant reported on the individual respondent's cognitive function. Data were collected using a computer-assisted personal interviewing system. Planned analyses include survey-weighted descriptive statistics and psychometric evaluation of the cognitive assessments. Ethics and dissemination Ethics approval was obtained from the ethics committees of the Institute of Nutrition of Central America and Panama, Maya Health Alliance, and the University of Michigan. Results will be disseminated through publications in peer-reviewed journals and presentations to local, national, and international audiences.

15
Bacterial metagenome in plaque, saliva, and tumor samples from individuals with and without OSCC by next-generation sequencing

ERIRA, A.; ROBAYO, D. A. G.; GAMBOA, F.; CHALA, A.; MORENO, A.; ARREGUI, A. C.; MUNOZ, E.; NOGUERA, J.; TOBAR-TOSSE, F.

2026-08-29 bioinformatics 10.64898/2026.08.27.747557 medRxiv
Top 2%
0.1%
Show abstract

Background: Oral dysbiosis has been associated with oral squamous cell carcinoma (OSCC); however, most microbiome studies rely on 16S ribosomal RNA (rRNA) gene sequencing, limiting species-level taxonomic resolution. Methods: Dental plaque, saliva, and tumor tissue samples from 10 patients with OSCC and dental plaque and saliva samples from 10 healthy controls were analyzed in this exploratory cross-sectional study. DNA was extracted and subjected to shotgun metagenomic sequencing using the Illumina MiSeq platform. Sequence reads were quality filtered with fastp, taxonomically classified using Kraken2 v2.1.3, and species-level abundances were re-estimated with Bracken v2.9 following the removal of human reads and low abundance taxa. Relative abundances were compared using the Mann Whitney U test with the Benjamini Hochberg false discovery rate correction, while the Bray Curtis principal coordinate analysis was used as an exploratory approach to visualize microbial community patterns. Results: Shotgun metagenomic sequencing revealed distinct bacterial community profiles across the oral microenvironment. Dental plaque exhibited the highest taxonomic diversity and relative abundance. The control plaque was enriched in Streptococcus koreensis, Capnocytophaga sp. oral taxon 878, Treponema sp. Marseille Q4132, and Leptotrichia sp. oral taxon 498, whereas the plaque from patients with OSCC showed a higher relative abundance of Pyramidobacter piscolens, Parvimonas parva, and Gemella sanguinis. Salivary samples displayed lower diversity and a more homogeneous composition, predominantly comprising Capnocytophaga endodontalis, Prevotella jejuni, Aggregatibacter aphrophilus, and Gemella sanguinis. The tumor tissue showed relatively higher abundance of Sellimonas catena, Escherichia coli, Solobacterium moorei, and Lacrimispora sp. HJ 01. Conclusions: This exploratory study provides species-level characterization of the oral microbiome across multiple oral microenvironments in OSCC and generates hypotheses for future integrative metagenomic and functional studies investigating the potential contribution of oral bacterial communities to OSCC pathogenesis.

16
Does genetic liability for autism influence alcohol use?

Page, S.; Easey, K.; Sedgewick, F.; Rai, D.; Stergiakouli, E.

2026-08-31 epidemiology 10.64898/2026.08.26.26360336 medRxiv
Top 2%
0.1%
Show abstract

A body of research suggests that autistic individuals are less likely to drink alcohol than neurotypicals. However, emerging studies support a link between autism and alcohol use. This complex relationship is also reflected in studies that have examined the genetic overlap between the two traits. However, it is unclear whether there is a direct causal relationship between them. To explore this, we applied a combination of polygenic score and Mendelian randomisation analyses using publicly available genome-wide summary statistics and phenotypic measures of autism and alcohol consumption from UK Biobank. LD score regression analyses did not provide evidence of a genetic correlation between genetic liability for autism and drinks consumed per week (rg=-0.08; CI95%=-0.19, 0.03). Further, findings from polygenic score analyses did not support an association between genetic liability for autism and overall monthly alcohol intake. Univariable Mendelian randomisation analyses showed little evidence for a total effect of autism, attention deficit hyperactivity disorder (ADHD) or depression on overall monthly alcohol consumption. Multivariable Mendelian randomisation analyses also showed little evidence of a direct effect of autism on drinks per week when controlling for ADHD and depression. It is plausible that genetic liability for autism does not directly increase the amount of alcohol consumed but instead operates via commonly co-occurring difficulties in the autistic community. However, our findings may be due to methodological shortcomings, including weak instruments biasing effects towards to the null. Consequently, results should be interpreted with caution and further research conducted to address these issues.

17
The impact of London's Ultra Low Emission Zone on respiratory prescribing: a synthetic control study

Williams, G. H.; Allen, T.

2026-09-01 epidemiology 10.64898/2026.08.27.26361515 medRxiv
Top 2%
0.1%
Show abstract

Urban air pollution remains a significant public health concern, contributing to premature deaths and adverse health outcomes. However, there is little causal research evaluating the effectiveness of policies designed to improve air quality. This study assesses the impact of all three stages of London's Ultra Low Emission Zone (ULEZ) on air pollution, via PM2.5 levels, and respiratory health, via prescription records for bronchodilator and respiratory corticosteroid medications. Analyses are at general practice level, using a generalised synthetic control method to estimate causal impacts. Stage 1 was associated with a statistically significant but negligible 0.77% reduction in PM2.5 levels, with no corresponding change in prescribing. Stage 2 produced a paradoxical 2.69% increase in PM2.5, alongside a 4.44% decrease in inhaled corticosteroid quantity but a 12.51% increase in average daily quantity (ADQ) usage, suggesting a worsening of disease severity among existing patients. Stage 3 yielded a 2.69% PM2.5 reduction and a modest 2.18% decrease in bronchodilator ADQ usage. Spillover effects beyond the ULEZ boundary were statistically significant, but negligible. We find overall that the ULEZ had minimal effects on both air quality and respiratory prescribing across all three stages. These findings provide new insights into the effectiveness of ULEZ policies in reducing air pollution and its associated health impacts, suggesting the zone's effects are considerably smaller than previously reported, and that integration with broader policy measures may be necessary to achieve meaningful public health gains.

18
Causal roles of phenotypic age and metabolic health on dementia: a Mendelian randomisation and structure learning study

Baousi, A.; Dobinda, K.; Zhu, J.; Yu, X.; Muir, K.; Lophatananon, A.; McMillan, B.; Clarkson, P.; Tang, E. Y. H.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.09.01.26360731 medRxiv
Top 3%
0.1%
Show abstract

Background Phenotypic age acceleration (PhenoAgeAccel), derived from PhenoAge, and MetaboHealth are composite exposures of biological ageing and metabolic health associated with dementia-related outcomes. Whether these associations are causal and reflect the exposures, constituent biomarkers, or both remains unclear. Methods This study included UK Biobank participants of White British genetic ancestry. MetaboHealth was derived from nuclear magnetic resonance (NMR) metabolomics and PhenoAgeAccel from clinical biomarkers and chronological age. Genome-wide association studies (GWAS) were conducted for MetaboHealth (n=272,568) and PhenoAgeAccel (n=274,077). Independent genome-wide significant variants were used as genetic instruments in two-sample Mendelian randomisation (MR) with FinnGen all-cause dementia summary statistics. Inverse-variance weighting was the primary MR method. Causal network analysis estimated relationships among constituent biomarkers and dementia. Findings GWAS identified 126 and 141 independent genome-wide significant variants for MetaboHealth and PhenoAgeAccel, of which 109 and 141 were retained as genetic instruments. MR found no evidence of a causal effect of genetically predicted MetaboHealth (per unit: OR 0.83, 95% CI 0.49-1.42; p=0.51) or PhenoAgeAccel (per year: OR 0.99, 95% CI 0.95-1.02; p=0.44) on all-cause dementia, with consistent findings across sensitivity analyses and robust MR methods. Lower lymphocyte percentage and higher NMR-derived glucose had direct relationships with dementia in the joint constituent-biomarker network. Interpretation MR provided no evidence that either composite exposure causally influenced dementia. The network prioritised lymphocyte percentage and NMR-derived glucose, supporting examination of composite exposures alongside their constituent biomarkers. Funding NIHR, UKRI, MRC, UK Dementia Research Institute, Innovate UK, and European Union. Full funding details are provided in the acknowledgements.

19
Relation of Self-Reported Race and Genetic Ancestry to Hypertension Prevalence Among Hispanics/Latinos: The Hispanic Community Health Study/Study of Latinos

Montanez-Valverde, R. A.; Kim, V.; Duran-Luciano, P.; Yuan, Y.; Sofer, T.; Kaplan, R. C.; Gallo, L. C.; Talavera, G. A.; Perreira, K. M.; Daviglus, M. L.; Rosas, S. E.; Llabre, M. M.; Elfassy, T.; Li, X.; Isasi, C. R.; Rodriguez, C. J.

2026-09-03 genetic and genomic medicine 10.64898/2026.09.01.26361995 medRxiv
Top 3%
0.1%
Show abstract

Background. The imprecision of current metrics to capture the complex genetic admixture and racial identity among Hispanic/Latino individuals in the United States [US] is a concern. We examined the relationship of self-reported race and genetic ancestry with hypertension [HTN] among Hispanics/Latinos. Methods. Cross-sectional study of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL), including 10,586 Hispanic/Latino unrelated adults. Genetic ancestry: West African [AA], Amerindian [AI], and European [EA]. Self-reported race: White, Black, Native American, or Multiple/Missing (More than one race or Unknown/Not reported/Refused). HTN: systolic (SBP) [&ge;]130 mmHg, diastolic blood pressure (DBP) [&ge;]80 mmHg, and/or use of HTN medications. Age- and sex adjusted models were used. Results. Self-reported race was White (38{middle dot}6%), Black (3{middle dot}6%), Native American (4{middle dot}1%), and Multiple/Missing (53{middle dot}7%), with Unknown/Not reported/Refused representing 32{middle dot}7%. Black and White Hispanics/Latinos had the greatest AA (55{middle dot}7%) and EA (69{middle dot}3%) ancestries, respectively. Each 10% AA increase was associated with OR 1{middle dot}15, SBP beta +0{middle dot}9 mmHg, and DBP beta +0{middle dot}7 mmHg. Conversely, each 10% AI increase was associated with OR 0{middle dot}83, SBP beta -0{middle dot}4 mmHg, and DBP beta -0{middle dot}6 mmHg. HTN prevalence was highest among those with Black race or in the highest AA quantile (45{middle dot}6% and 48{middle dot}0%, respectively), and lowest among those with Native American race or in the highest AI quantile (37{middle dot}6% and 26{middle dot}7%, respectively). Conclusion. One-third of Hispanics/Latinos did not self-report race. Black or White self-reporting race did somewhat relate to AA or EA ancestry, respectively. HTN profiles were related to self-reported race and genetic ancestry in this admixed population.

20
A Curated Pharmacogenomic Allele Catalog for Sub-Saharan African Populations

SULAIMAN, M. A.; Oyeyemi, B. F.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.25.26361354 medRxiv
Top 3%
0.1%
Show abstract

Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.